Papers with multi-armed bandit learning

1 papers
Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Recent research on test-time adaptation suggests a possible way to improve the generalization ability of LLMs.
Approach: They propose to use multi-armed bandit learning and multi-arm dueling bandits to solve a multi-source test-time model adaptation problem from user feedback.
Outcome: The proposed model is more effective than other strong baselines on extractive question answering datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations